Papers by Thomas L. Griffiths
Evaluating distillation methods for data-efficient syntax learning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | knowledge distillation (KD) targeting attention should selectively accelerate syntax acquisition, a study finds . logit-based KD dramatically improves data-efficiency, attention-based one provides minimal benefit even for syntactic tasks. |
| Approach: | a study predicts that knowledge distillation targeting attention should selectively accelerate syntax acquisition . a systolic analysis of student models compared to logit-based knowledge distillations . |
| Outcome: | a new study shows that knowledge distillation (KD) targeting attention accelerates syntax acquisition . the hypothesis is tested on syntactic benchmarks and perplexity. |
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation (2026.findings-acl)
Copied to clipboard
| Challenge: | Reinforcement Learning from Hindsight Simulation (RLHF) can cause severe misalignment in generative AI, but it is not a universal method for fine-tuning large language models. |
| Approach: | They propose a method that uses evaluator feedback to decouple alignment signal from potentially compromised predictions. |
| Outcome: | The proposed method significantly outperforms RLHF in comparisons with baselines and human evaluations. |
Localized Cultural Knowledge is Conserved and Controllable in Large Language Models (2026.findings-acl)
Copied to clipboard
Veniamin Veselovsky, Berke Argın, Benedikt Stroebl, Chris Wendler, Robert West, James Evans, Thomas L. Griffiths, Arvind Narayanan
| Challenge: | Large language models (LLMs) display language patterns influenced by their native tongue when learning new languages. |
| Approach: | They propose to quantify the explicit-implicit localization gap in large language models by using a new cultural localization benchmark and find large gaps in the majority of models. |
| Outcome: | The proposed model can generate culturally localized responses in multiple languages while maintaining language accuracy and task diversity. |